Skip to content

sleep: measure the remaining time against a clock, not the kernel's remainder - #642

Merged
xrme merged 1 commit into
Clozure:masterfrom
cyclistmass:sleep-clock-remainder
Sep 29, 2026
Merged

xrme merged 1 commit into
Clozure:masterfrom
cyclistmass:sleep-clock-remainder

Conversation

@cyclistmass

Copy link
Copy Markdown
Contributor

(sleep n) overruns when another thread allocates, and the error grows with the collection rate. This is issue #639.

%nanosleep sleeps again for the remaining time the kernel writes back on EINTR. The kernel computes that remainder before it runs the signal handler, and suspend_resume_handler blocks inside the handler until the collection finishes, so the time the thread spends parked is never subtracted. Every world stop adds its own duration to the sleep.

There is one definition, under #-windows-target, with no per-port override, so every non-Windows target has this. Three callers reach it: sleep, the periodic-task sleep in housekeeping-loop, and process-wait once per poll tick.

The patch works the deadline out once, from a clock the Lisp reads itself, then sleeps again for the remaining time it measures on each EINTR. It never reads the kernel write-back, so #_nanosleep now gets a NULL rem, and the Darwin negative-remainder workaround goes with the code it guarded.

That gives one loop for every target. The clock is CLOCK_MONOTONIC through #_clock_gettime, which lib/time.lisp already reads the same way, and gettimeofday on Darwin, a kernel import that needs no interface-database entry. I avoided clock_nanosleep with TIMER_ABSTIME because Darwin does not have it, which would mean a second loop with a different error convention for one target.

Measured

Red then green on both architectures at 6526e21c, changing only this patch, with extended-tests/threads/sleep-vs-alloc.lisp from Clozure/ccl-tests#10.

arch without the patch with it
linuxarm64 170.1 s for a 10 s sleep, ratio 17.0 10.0 s, ratio 1.0
linuxx8664 31.8 s, ratio 3.2 10.0 s, ratio 1.0

On arm64 the kernel md5 is identical in both halves, as it should be, since the patch touches no kernel file. Both machines are 2-vCPU burstable instances, so read those as verdicts and not as benchmark figures.

ANSI reports 21679/0 and ccl-tests 243/0 with the patch applied.

I have not tested Darwin. There is no Darwin build here, so the gettimeofday arm rests on construction rather than measurement.

@xrme

xrme commented Sep 24, 2026

Copy link
Copy Markdown
Member

Since you opened the PR, master has gained a monotonic clock. In 624854d, I added a new kernel import lisp_monotonic_time, and 1e88afc uses that to implement get-internal-real-time. So, get-internal-real-time is now monotonic on every platform: CLOCK_BOOTTIME on Linux, CLOCK_MONOTONIC_RAW on Darwin, QueryPerformanceCounter on Windows, and CLOCK_MONOTONIC on FreeBSD and illumos.

Can you rebase and use the improved get-internal-real-time for the deadline instead of %nanosleep-clock? This would match what a similar deadline loop does in %timed-wait-on-semaphore-ptr.

Some data from other platforms, from master without this PR's changes (ccl --load sleep-vs-alloc.lisp):

platform result
darwin SLEEP-VS-ALLOC-RESULT :REQUESTED 10 :ELAPSED 10.0 :RATIO 1.0 :ALLOCATIONS 315351 :VERDICT PASS
freebsd SLEEP-VS-ALLOC-RESULT :REQUESTED 10 :ELAPSED 26.0 :RATIO 2.6 :ALLOCATIONS 1057444 :VERDICT FAIL
illumos SLEEP-VS-ALLOC-RESULT :REQUESTED 10 :ELAPSED 26.9 :RATIO 2.7 :ALLOCATIONS 619815 :VERDICT FAIL

Darwin's nanosleep apparently works out the remaining time after the signal handler returns, so the time spent parked is counted. FreeBSD and illumos behave like Linux.

…emainder

(sleep n) returns late, and on a machine that collects often enough it does
not return at all, when another thread allocates.  This is issue Clozure#639.

%nanosleep calls #_nanosleep and, when a signal interrupts the call, sleeps
again for the remaining time the kernel wrote back.  Each GC suspends the
sleeping thread with a signal, and suspend_resume_handler blocks inside the
handler until the collection is over.  The kernel computes the remaining time
before it runs the handler, so the time the thread spends blocked in the
handler is not subtracted.  Every world stop adds its own duration to the
sleep, and the error grows with the collection rate.  On two 2-vCPU burstable
instances a 10 s sleep took 31.8 s on linuxx8664 and 170.1 s on linuxarm64.
Earlier runs on the same arm64 cell were killed at 90 s without returning, so
the overrun has no bound that I have established.

Darwin does not have the defect: its nanosleep works the remaining time out
after the signal handler returns, so the parked time is counted.  FreeBSD and
illumos behave like Linux.  Those three measurements are not mine; they come
from the maintainer, who ran the reproducer on platforms I cannot reach.

The fix computes the deadline once and, on each EINTR, sleeps again for
(stop - now).  The kernel's write-back is no longer read at all, so
#_nanosleep now gets a NULL rem pointer, and the Darwin check for a negative
zero-extended remainder goes with the code it guarded.

The clock is get-internal-real-time, which is monotonic on every platform
since 1e88afc.  The loop is shaped like the one in
%timed-wait-on-semaphore-ptr, which solves the same problem for a semaphore
wait: keep a stop in internal-time-units, re-read the clock on each wakeup,
and floor the difference back into the units the call wants.
clock_nanosleep with TIMER_ABSTIME was considered and not used: Darwin does
not have it, so it would need a second loop with a different error convention
for one target.  This form is one loop for every target.

Cost: two clock reads per %nanosleep call and one more per interruption.
process-wait polls through %nanosleep once per tick, so this is noise there.

Measured red then green with ccl-tests extended-tests/threads/sleep-vs-alloc.lisp
on linuxx8664 and linuxarm64.  The pull request carries the cell.  Not built on
a 32-bit port, and not built on Darwin, FreeBSD or illumos.
@cyclistmass

Copy link
Copy Markdown
Contributor Author

Rebased onto master and rewritten to use get-internal-real-time, as you asked.

%nanosleep-clock is gone. The loop now follows %timed-wait-on-semaphore-ptr.
It binds now and stop up front. On EINTR it re-reads the clock and returns
once now has reached stop. Otherwise it rebuilds the remaining time with
floor against internal-time-units-per-second.

Two things fall out of that shape.

The call passes (%null-ptr) for the remainder, because nothing reads it any
more. That also retires the #+(and darwin-target 64-bit-target) guard for the
negative-remainder bug. The guard only ever detected a bad value in the kernel's
remainder. No platform consults the remainder now.

The two-timespec ping-pong goes with it. One timespec is refilled from the
recomputed seconds and nanoseconds on each pass.

@xrme
xrme merged commit b306852 into Clozure:master Sep 29, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants